Papers with semantic analysis

18 papers
NLP Workbench: Efficient and Extensible Integration of State-of-the-art Text Mining Tools (2023.eacl-demo)

Copied to clipboard

Challenge: NLP Workbench is a web-based text mining platform that allows non-expert users to obtain semantic understanding of large-scale corpora using state-of-the-art text mining models.
Approach: They propose to use a microservice architecture to replace existing models or integrate a new one.
Outcome: The proposed model is extensible and can be easily replaced or integrated with existing models.
What does the language of foods say about us? (D19-62)

Copied to clipboard

Challenge: Using a dataset of 24 million food-related tweets, we can predict if states in the United States are above the median rates for type 2 diabetes mellitus (T2DM) income, poverty, and education are important factors in predicting T2DM rates, but socioeconomic factors do not capture this information.
Approach: They use a dataset of 24 million food-related tweets to investigate the signal contained in the language of food on social media.
Outcome: The language of food can predict health risks, political orientation, and geographic location, and outperform previous work by 4–18%.
SKICSE: Sentence Knowable Information Prompted by LLMs Improves Contrastive Sentence Embeddings (2024.naacl-short)

Copied to clipboard

Challenge: Experimental results show that our model outperforms the previous state-of-the-art model GloVe on STS tasks.
Approach: They propose a simple and effective prompt template that is able to obtain the knowable information of input sentences from LLMs.
Outcome: The proposed model outperforms the previous state-of-the-art model PromptBERT on STS tasks.
CASE: Large Scale Topic Exploitation for Decision Support Systems (2025.coling-demos)

Copied to clipboard

Challenge: Topic models are still a major tool for information retrieval and summarization, but their integration into decision-making systems is limited.
Approach: They propose a tool for exploiting topic information for semantic analysis of large corpora using a Solr engine and a customized indexing strategy.
Outcome: The proposed approach can be used to analyze large corpora and perform thematic trend analysis, topic-based document retrieval, or similarity search.
TEACH: A Contrastive Knowledge Adaptive Distillation Framework for Classical Chinese Understanding (2025.acl-long)

Copied to clipboard

Challenge: Traditional methods for processing classical Chinese segment language understanding into discrete tasks, which overlook crucial background information and reduce user engagement.
Approach: They propose a framework that integrates word sense disambiguation with sentence translation to minimize hallucinations and improve semantic analysis.
Outcome: The proposed framework integrates word sense disambiguation with sentence translation to minimize hallucinations and improve semantic analysis.
Universal Proposition Bank 2.0 (2022.lrec-1)

Copied to clipboard

Challenge: Semantic role labeling (SRL) is a shallow semantic parsing task that identifies "who did what to whom when, where etc." SRL is useful in a wide range of downstream NLP tasks and real-world applications.
Approach: They propose a method to generate shallow semantic parsing tasks using monolingual SRL and multilingual parallel data.
Outcome: The proposed method improves the quality of the generated propbanks.
Cross-lingual Decompositional Semantic Parsing (D18-1)

Copied to clipboard

Challenge: Renewed interest in semantic analysis has led to a surge of proposed new frameworks . many of these efforts are limited to the analysis of English, but with a number of exceptions e.g., recent efforts in Minimal Recursion Semantics (MRS) and multilingual FrameNet annotation and parsing.
Approach: They propose a cross-lingual decompositional semantic analysis task based on a target language . they propose 'end-to-end' model with an annotating mechanism that supports intra-sentential coreference .
Outcome: The proposed model outperforms baselines by at least 1.75 F1 score on an evaluation dataset.
CodeQA: A Question Answering Dataset for Source Code Comprehension (2021.findings-emnlp)

Copied to clipboard

Challenge: False. a free-form question answering dataset can serve as a useful research benchmark for source code comprehension.
Approach: They propose a free-form question answering dataset for source code comprehension . they implement syntactic rules and semantic analysis to transform code comments into question-answer pairs.
Outcome: The proposed dataset can serve as a useful research benchmark for source code comprehension.
Evaluating Scoped Meaning Representations (L18-1)

Copied to clipboard

Challenge: Semantic parsing offers many opportunities to improve natural language understanding . current research on open-domain semantic parsers focuses on supervised learning methods .
Approach: They propose a semantically annotated parallel corpus for English, German, Italian, and Dutch . they use a matching tool to evaluate scoped meaning representations to match clauses .
Outcome: The proposed method captures the semantics of negation, modals, quantification, and presupposition triggers . it compares scoped meaning representations to gold standard parsers and finds improvements .
Construct a Sense-Frame Aligned Predicate Lexicon for Chinese AMR Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing lexicons blur senses and frames of predicates, which needs to be refined to meet word sense disambiguation and event extraction tasks.
Approach: They propose to construct a predicate lexicon for Chinese AMR corpus with 14,389 senses and 10,800 frames of 8,470 words.
Outcome: The proposed lexicon includes 14,389 senses and 10,800 frames of 8,470 words.
Towards the Inference of Semantic Relations in Complex Nominals: a Pilot Study (L18-1)

Copied to clipboard

Challenge: Complex nominals (CNs) show similar external forms but encode different semantic relations because of noun packing.
Approach: They propose to use paraphrases to convey conceptual content of english two-term CNs in the domain of environmental science to disambiguate the semantic relation between constituents of CN.
Outcome: The proposed method disambiguates the semantic relation between constituents of the CN and infers the semantic relations in these multi-word terms.
A Semantic-Aware Layer-Freezing Approach to Computation-Efficient Fine-Tuning of Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Existing work on how to finetune but neglects the issue of where to fine-tune language models is expensive.
Approach: They propose to use transition traces of latent representation to compute deviations (or loss) and then estimate the gain of each layer in reducing deviation (or gain).
Outcome: The proposed approach outperforms baseline methods and is cost-benefit balanced.
Efficient AMR Parsing with CLAP: Compact Linearization with an Adaptable Parser (2024.lrec-main)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) parsers face efficiency challenges because of their large model size and computational time, which limit their accessibility within the research community.
Approach: They propose a novel linearization system that simplifies encoding and reduces the number of tokens by between 40% and 50%.
Outcome: The proposed system reduces the number of tokens by 40% and 50% while maintaining high performance while reducing training and inference times.
RoBERT2VecTM: A Novel Approach for Topic Extraction in Islamic Studies (2024.findings-emnlp)

Copied to clipboard

Challenge: a new approach to investigate “Hadith” texts presents challenges due to the complexity of Arabic . a novel neural-based approach to analyze “Matn” topics outperforms traditional NLP models .
Approach: They propose a novel approach to analyze Arabic “Hadith” texts using the Contextualized Topic Model.
Outcome: The proposed approach outperforms state-of-the-art models by generating more coherent topics in Arabic.
On the Role of Semantic Proto-roles in Semantic Analysis: What do LLMs know about agency? (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies on large language models (LLMs) have not explored their capacity to reason over event structure . et al., 2015, 142: e007-e0027; eugene, 1985; Weiner, 1995; saab, 1985) focus on the role of large language model in decision-making .
Approach: They propose to characterize agents via properties such as "instigation" and "volition" they also examine whether incorporating semantic proto-role labeling context improves SRL performance .
Outcome: The proposed model improves in a zero-shot setting by incorporating proto-role labeling context . the results support previous work showing that LLMs underperform human annotators in complex semantic analysis.
Survey on Thai NLP Language Resources and Tools (2022.lrec-1)

Copied to clipboard

Challenge: Thai language is one of the under-resourced languages in the NLP domain, although it is spoken by approximately 70 million people globally.
Approach: They propose to use Thai language as an example to understand how NLP works and how it can be applied to Thai language.
Outcome: The results show that Thai NLP research has progressed over the past three decades, especially on upstream tasks such as tokenisation, but research on downstream tasks such syntactic parsing and semantic analysis is still limited.
Towards Optimal Evaluation Efficiency for Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) require large-scale benchmarks, which are costly in terms of time, computational resources, or API tokens.
Approach: They propose an efficient evaluation framework that selects a question subset based on pre-tested results and uses semantic analysis to evaluate whether the subset preserves the original benchmark.
Outcome: The proposed evaluation framework outperforms previous methods in reliability and score accuracy.
Reasoning Hijacking: The Fragility of Reasoning Alignment in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Current LLM safety research focuses on mitigating **Goal Hijacking**, preventing attackers from redirecting a model’s high-level objective.
Approach: They propose a new adversarial prompt attack paradigm that subverts model judgments by injecting spurious decision criteria without altering the high-level task goal.
Outcome: The proposed model subverts model judgments by injecting spurious decision criteria without altering the high-level task goal.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations